Back

Clinical Trials

SAGE Publications

Preprints posted in the last 30 days, ranked by how well they match Clinical Trials's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
A Measurement-Based Care Strategy for Buprenorphine-Naloxone Treatment (Bup-MBC): Development of an EHR-Integrated Intervention

Reese, T.; Audet, C.; Ancker, J.; Wright, A.; Marcovitz, D.; Kast, K. A.; Bridges, J.; Tindle, H.; Shah, M.; von Horn, A.; Matheny, M. E.

2026-09-01 addiction medicine 10.64898/2026.08.27.26361539 medRxiv
Top 0.1%
18.6%
Show abstract

Introduction: Risk of recurrent opioid use during buprenorphine-naloxone (bup-nx) treatment is dynamic and remains elevated after initiation, with vulnerability shaped in part by treatment intensity and gaps between visits, yet routine outpatient care relies on episodic encounters and retrospective data. This mismatch can delay recognition of emerging instability and limit timely treatment adjustments. This paper reports the development and specification of an intervention strategy to address this mismatch. Methods: We used a structured, multi-phase design process to specify and configure a measurement-based care (MBC) strategy for bup-nx treatment (Bup-MBC) in outpatient addiction clinics through three phases: (1) a systematic review of patient-reported outcome measures (PROMs) for substance use treatment; (2) a qualitative needs assessment using the Theoretical Domains Framework and COM-B (Capability, Opportunity, Motivation-Behavior) model to identify gaps in risk monitoring, agency, and trust; and (3) iterative co-design with multidisciplinary clinicians to refine workflow fit and trust-preserving use of data. Patients informed item and feedback content during the needs assessment but did not participate in the co-design cycles. Results: Bup-MBC integrates (1) brief between-visit PROMs (e.g., withdrawal, craving, adherence); (2) immediate non-punitive patient feedback; (3) clinician-facing summaries and non-directive prompts in the electronic health record (EHR); and (4) an opt-in between-visit outreach pathway with predefined safety triggers, all configured within existing EHR and patient portal infrastructure. It targets patient and clinician capability to recognize changes in risk, opportunity for action through structured monitoring and visit preparation, and trust and agency through non-punitive communication, without adding substantial burden. The full measure set, severity bands, and question-to-action map are provided as supplementary material. Key trade-offs included prioritizing single-item measures for feasibility, balancing opt-in outreach with safety overrides, and assuming routine clinician use of summaries. Conclusion: This development study specifies an EHR-integrated MBC strategy for outpatient bup-nx treatment. As single-center design work with co-design limited to clinicians and delivery contingent on portal or text-message access, its outputs are hypotheses about mechanism and fit rather than demonstrated effects. Feasibility studies are needed to evaluate uptake, acceptability, workflow fit, and effects on treatment.

2
Comparative Effectiveness of Single vs. Dual WhatsApp Reminders on No-shows: A Target Trial Emulation within the Public Health System of Buenos Aires, Argentina.

Esteban, S.; Quintana, G.; Sanchez, M.; Szmulewicz, A.

2026-08-19 health systems and quality improvement 10.64898/2026.08.17.26360609 medRxiv
Top 0.1%
13.1%
Show abstract

Background: Digital reminders reduce outpatient no-shows, but the optimal timing and frequency of messages remain unclear, particularly in Latin American public health systems. We emulated a target trial to evaluate the comparative effectiveness of four WhatsApp reminder strategies on appointment absenteeism and patient-initiated cancellations. Methods: We analyzed administrative and electronic health-record data from the public health system of the Autonomous City of Buenos Aires, Argentina (June 2023-May 2024). Eligible individuals had scheduled an in-person outpatient appointment in one of 15 prioritized specialties at least 75 hours in advance and had a mobile phone on record. We compared four strategies: (1) dual reminders at ~72 and ~24 hours before the appointment; (2) a single reminder at ~72 hours; (3) a single reminder at ~24 hours; and (4) no reminders. The primary outcome was the proportion of no-shows by the end of follow-up. Secondary outcomes were the cumulative incidence of patient-initiated cancellations overall, within 12 hours of the appointment, and followed by rebooking. We emulated the target trial using a cloning-censoring-weighting approach to estimate per-protocol controlled direct effects, with inverse-probability weights to address time-varying confounding and selection bias. Cumulative incidence of secondary outcomes was estimated using weighted Kaplan-Meier curves. Three pre-specified sensitivity analyses and standardized mean differences assessed robustness and covariate balance. Results: A total of 475,214 first eligible person-appointments were included; baseline no-show risk in the control arm was 34.6%. All three active strategies reduced no-shows compared with no reminders. The single 24-hour reminder produced the largest reduction (Risk Ratio [RR] 0.76, 95% CI 0.72, 0.81; Risk Difference [RD] -8.21 percentage points [pp], 95% CI -9.68, -6.54), followed by the dual-reminder strategy (RR 0.80, 95% CI 0.79,0.81; RD -7.05 pp, 95% CI -7.41, -6.71) and the single 72-hour reminder (RR 0.91, 95% CI 0.84,0.99; RD -3.16 pp, 95% CI -5.69, -0.49). All active strategies increased patient-initiated cancellations relative to control, with the dual-reminder strategy producing the largest increase. Sensitivity analyses preserved the qualitative ranking of strategies across all specifications. Conclusions: In this large target trial emulation, a single just-in-time WhatsApp reminder sent ~24 hours before the appointment was as effective as a dual-reminder schedule in preventing no-shows and superior to a distal 72-hour reminder alone. Adding a second, distal reminder provided no measurable benefit for attendance but substantially increased patient-initiated cancellations, which may be operationally valuable when active slot reallocation is a goal. These findings support timing, rather than frequency, as the primary lever of digital-reminder effectiveness, and favor the deployment of a single proximal reminder as the default strategy in resource-constrained outpatient settings.

3
Insights from a double-blind, randomized, direct-to-participant intervention trial for Long COVID

Vogel, J. M.; Ter Meer, J.; Foster-Bonds, R.; Duff, M. P.; Goosen, A.; Kurakova, A.; Dinh-Luong, E.; Miyasaki, L.; Topol, S.; Sturm, C.; Nowak, C.; Tate, A.; Redd, J.; Shepard, C.; Kheterpal, V.; Steinhubl, S. R.; Topol, E. J.

2026-08-22 infectious diseases 10.64898/2026.08.19.26360832 medRxiv
Top 0.1%
9.7%
Show abstract

Background. Long COVID affects an estimated 400 million people worldwide, and is associated with low quality of life. Nearly all completed Long COVID clinical trials reported no benefit, and most required participants to travel to study sites. This requirement systematically excludes severely affected patients. Because there are numerous candidate therapeutics with established safety profiles and regulatory approvals for other indications, scaled, efficient evaluation of therapeutics is needed. Methods. We designed and are conducting a double-blind, placebo-controlled, phase two trial of tirzepatide for Long COVID fatigue, using an entirely remote infrastructure. Design elements included electronic consent, identity and diagnosis verification through document upload, cold-chain delivery of an injectable study drug through a central pharmacy, shared decision-making for dose titration, repeated at-home capillary blood collection in a biospecimen subcohort, weekly participant touch points through study application, wrist-worn wearable monitoring, and clinical support. The trial is operating under FDA Investigational New Drug authorization. Results. This trial enrolled 1,058 participants in 73 days, at least double the rate of any other Long COVID trial. Mean baseline metrics include mean Fatigue Severity Scale of 59.3 (standard deviation [SD] 4.9), daily step count of 3,611 (SD 2,706, general population reference mean 7,731), EQ-5D-5L of 0.6 (SD 0.2), and FUNCAP27 4.0 (SD 1.0), which was a more severely affected population than other clinical trials that collected comparable data. Study processes are working as designed. Participants use existing advocacy and support channels to gather and communicate. Conclusions. A direct-to-participant, siteless infrastructure can support a double-blind placebo-controlled trial of an injectable drug at scale, accelerate accrual, and reach severely affected participants who are routinely excluded by site-based designs. Modernizing drug distribution and regulatory pathways is needed to realize the full potential of decentralized infrastructure for drug repurposing clinical trials.

4
New tests for trials of very few patients using longitudinal data - a case-study in Autosomal Recessive Cerebellar Ataxias

Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.

2026-09-02 health informatics 10.64898/2026.08.28.26361588 medRxiv
Top 0.1%
6.6%
Show abstract

We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.

5
Towards understanding the disease landscape of clinical trials in Germany: Ontology and embedding-based pipelines versus Large Language Models for ICD-10 Harmonization

Ndabashinze, R.; Franzen, D.; Kozuch, E.; Aagerup, J.; Fink, A.; Yerunkar, S. S.; Hunter, K.; Mayo-Wilson, E.; Ying, X.; Kilicoglu, H.; Schorr, S. G.; Seidler, A. L.

2026-08-06 health informatics 10.64898/2026.08.04.26359616 medRxiv
Top 0.1%
3.9%
Show abstract

Background Clinical trials conducted in Germany are registered across multiple registries, including the German Clinical Trials Register (DRKS), ClinicalTrials.gov, the EU Clinical Trials Register (EUCTR), and, since 2023, the Clinical Trials Information System (CTIS). These registries record health conditions using different classification systems and terminologies, including ICD-10-GM, MeSH, MedDRA, and free text, making cross-registry analyses difficult. We developed and evaluated a pipeline for harmonizing trial condition descriptions to WHO ICD-10 and compared its performance with that of a large language model (LLM) and to health conditions coded by humans. Methods We developed a four-stage, registry-aware mapping pipeline consisting of: (i) condition mention extraction and normalization; (ii) classification of ICD-mappable versus non-mappable mentions; (iii) ontology-based candidate generation using UMLS links between MeSH, MedDRA, ICD-10-GM, and WHO ICD-10; and (iv) SapBERT-based semantic retrieval with hybrid confidence scoring. A second variant additionally applied cross-encoder reranking of the top candidate codes. A stratified sample of 500 condition mentions was manually coded to create an expert reference standard. GPT-4o was evaluated in parallel using the same structured decision framework as the human reviewers. Performance was assessed using accuracy, precision, F1 score, and Cohen's {kappa} at the three-character, block, and chapter levels of ICD-10. Results The pipeline was applied to 23,061 clinical trials and identified 39,512 ICD-mappable condition mentions, of which 72.4% received a high-confidence assignment. Against 390 expert-coded mentions, the baseline pipeline achieved 49.0% accuracy at the three-character ICD-10 level ({kappa} = 0.487), increasing to 58.7% at the chapter level ({kappa} = 0.561). The cross-encoder method produced small but consistent improvements across all evaluation levels. Candidate-recall analysis showed that the correct code was present in the retrieved candidate set in only 73.7% of cases. The LLM substantially outperformed both pipeline variants, achieving 96.7% accuracy and near-perfect agreement with expert coding ({kappa} = 0.966) at the three-character level. The LLM also assigned clinically plausible codes to 82.4% of rejected mentions, 62.8% of Tier-3 exclusions, and 92.3% of review-band mentions. Conclusion Automated harmonization of clinical trial condition data across heterogeneous registries is feasible and supports the use of a common ICD-10 framework for cross-registry analyses. The LLMs achieved high agreement with expert coding, and performed better than the deterministic ontology and embedding pipeline, which achieved moderate agreement. These findings indicate that LLMs can support analyses of the distribution of health conditions investigated in clinical trials in Germany.They are a promising tool for classification of other non-standardised trial characteristics in registries. Keywords: Clinical trial registries; ICD-10; disease harmonization; UMLS; entity linking; SapBERT; large language models; clinical research; natural language processing.

6
Balancing Relapse Risk and Agency in Buprenorphine-naloxone Treatment: A Qualitative Needs Assessment to Inform Patient-Centered Care

Reese, T.; Shah, M. V.; Wright, A.; Matheny, M. E.; Marcovitz, D. E.; Kast, K. A.; Bridges, J.; Tindle, H.; von Horn, A.; Audet, C.

2026-08-23 addiction medicine 10.64898/2026.08.21.26360804 medRxiv
Top 0.1%
3.7%
Show abstract

Objectives Outpatient buprenorphine-naltrexone (bup-nx) treatment reduces overdose risk, yet many patients still return to use or disengage from treatment. We sought to understand how patients and prescribers experience and manage relapse risk, monitoring, and treatment agency in routine bup-nx treatment to identify gaps in current practice. Methods We conducted a qualitative needs assessment using semi structured, critical incident interviews with patients receiving outpatient bup-nx and prescribers who manage bup-nx treatment. Interviews examined situations involving relapse risk and empowerment in treatment decisions. We structured data collection and analysis using the Theoretical Domains Framework and COM B model to characterize determinants. Transcripts were coded deductively and inductively until code level saturation was reached. Results Participants (9 patients, 8 prescribers) described nine treatment needs mapped to the Capability, Opportunity, and Motivation components of the COM B model. These themes highlighted how patient agency in bup-nx treatment was constrained by physiologic and emotional states, with withdrawal, craving, pain, and distress often overriding longer term goals. Relapse vulnerability was experienced as dynamic and intensifying between visits, while clinical detection remained anchored to visit bound assessments, urine drug testing, refill patterns, and crisis driven contact, creating blind spots. Structural friction (pharmacy rules, insurance disruptions, transportation and housing instability), stigma from family and recovery communities, and motivational processes tied to fluctuating readiness and trust in monitoring further shaped engagement, disclosure, and dosing decisions; the same monitoring tools could either support honest disclosure or provoke concealment when perceived as punitive. Conclusions Relapse risk and agency in bup-nx treatment are negotiated as dynamic processes within structurally constrained and trust sensitive systems. Addressing the identified capability, opportunity, and motivation gaps will require patient centered, trust preserving approaches to monitoring and shared decision making.

7
Bayesian Borrowing of External Information in Clinical Trials: A Comparison of MAP, RMAP, and SAM Priors

Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.26.26360843 medRxiv
Top 0.1%
3.3%
Show abstract

Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.

8
From Output Errors to Workflow Harm: A Practitioner-Audit Method for LLM-Mediated Research

Austria, D.; McCollister, B.; Lindsey, J. E.; Arowolo, M.; Okon, M.

2026-08-17 health informatics 10.64898/2026.08.13.26360414 medRxiv
Top 0.1%
3.2%
Show abstract

Objective. Formal large language model (LLM) evaluations score isolated prompts, but clinicians and health-informatics researchers meet model failures inside multi-step workflows where erroneous output can alter procedures or contaminate documents. We present TRACE (Tracking Reliability of AI-generated Conversational Evidence), a practitioner-audit framework for evaluating the downstream workflow reliability of conversational AI. Materials and Methods. A method paper with an empirical demonstration: 45 documentation-positive incidents recorded by one clinician-informatician across scholarly, clinical informatics, and clinical-adjacent workflows over seven weeks, coded with a consequence-based severity rubric, an error definition, a taxonomy crosswalk, and a Response-Audit Scorecard. Three reviewer-authors independently coded a 16-incident subsample; three vendor-blinded AI comparators applied the taxonomy to all 45 incidents. Results. Four categories tied as most frequent: verification failure, factual numerical error, tool-behavior misunderstanding, and citation or reference formatting (n=7 each). Four workflow-harm patterns recurred: procedural propagation, documentary contamination, trust-calibration disruption, and user-borne corrective burden, and one incident carried an estimated $2500 impact. Category agreement across three human reviewer-authors was low (Fleiss {kappa}=0.155), whereas three AI comparators agreed substantially (Fleiss {kappa}=0.632), suggesting taxonomy legibility under standardized conditions even where human judgment diverged. Discussion. Category assignment is comparatively legible, whereas severity and claimed-verification remain judgment-dependent. The claimed-verification gap is a measurable failure mode distinct from hallucination, sycophancy, and over-refusal. Conclusion. Practitioner audits with structured response scoring complement benchmarks by documenting workflow harm as an applied evaluation unit for clinical informatics and public-health work; this is a pilot that motivates, not estimates, error rates or cross-model comparisons.

9
Adverse Drug Events Across Data-Production Contexts: Multilingual Detection, Alignment, and Cross-Genre Discourse Analysis

Ma, Y.; Weissenbacher, D.; Patock, J.; Gonzalez-Hernandez, G.

2026-08-11 health informatics 10.64898/2026.08.08.26360012 medRxiv
Top 0.1%
2.6%
Show abstract

Adverse drug event (ADE) evidence is produced across patient-generated, clinical, and scientific settings that differ in language, documentation purpose, terminology, and degree of standardization. These differences shape both which adverse experiences become visible to pharmacovigilance systems and how readily they can be linked to curated drug-safety knowledge. We examine these relationships across five corpora representing distinct data-production settings: ADE Corpus V2 (medical case reports), SMM4H-2026 Task 1 (multi-lingual user-generated health content), CADEC V2 (patient-forum narratives), the Dutch ADE Corpus (EHR clinical notes), and TwiMed-PubMed (biomedical literature). A shared BERTopic analysis of ADE-positive texts concerning antidepressants and antihypertensives across the four English-language corpora identified nine interpretable topics. CADEC V2 contained a more differentiated distribution of symptom-specific themes, including sexual effects, suicidal or panic-related thoughts, vivid dreams, and memory difficulties, whereas SMM4H-2026, TwiMed-PubMed, and ADE Corpus V2 were dominated by a broader medication, sleep, tiredness, and pain theme. These patterns indicate that data-production context shapes what adverse experiences are expressed and standardized, with patient-generated narratives surfacing subjective, symptom-specific experience largely absent from clinical and scientific sources. We further show that this context shapes how readily real-world drug mentions can be linked to curated pharmacovigilance knowledge. Using SIDER 4.1 as a retrieval resource, we find substantial cross-corpus mismatches between real-world drug mentions and SIDER's predominantly English, generic-name vocabulary: CADEC V2 achieved only 9.5% exact-match coverage, with unmatched mentions frequently involving brand names, misspellings, and language-specific variants, compared to 91.0% coverage in TwiMed-PubMed's formally standardized biomedical literature. To probe how these representational differences interact with automated detection, we compare corpus-specific QLoRA fine-tuning of Llama-3.2-3B with retrieval-augmented inference using Llama-3.1-70B and Llama-3.1-405B grounded in SIDER-retrieved evidence. QLoRA-Llama-3B achieved the highest micro-averaged F1 scores on ADE Corpus V2 (0.91), CADEC V2 (0.88), and SMM4H-2026 (0.80), whereas SIDER-grounded inference with Llama-3.1-405B achieved the highest scores on Dutch ADE (0.95) and TwiMed-PubMed (0.91); these corpus-dependent patterns should not be interpreted as a controlled comparison of adaptation strategies, since model scale, task formulation, and available supervision differ across datasets. Together, our findings indicate that data-production context influences what adverse experiences are expressed, how they are standardized, and how readily they can be retrieved and computationally detected. Pharmacovigilance systems should therefore combine source-sensitive supervision with external knowledge grounding while explicitly monitoring gaps between real-world language and curated drug-safety resources.

10
Effect of transitioning virally suppressed children and adolescents with HIV to dolutegravir-based antiretroviral therapy: emulated target trials in a large cohort in South Africa

Brown, J. A.; Sookrajh, Y.; Mtila, L.; Lushaba, N.; Hlabisa, M.; van der Molen, J. S.; Tlhaku, K.; Nkosi, M.; Ngwenya, T.; Khubone, T.; Mahomed, S.; Chammartin, F.; Archary, M.; Garrett, N.; Lewis, L.; Dorward, J.

2026-08-22 hiv aids 10.64898/2026.08.19.26360677 medRxiv
Top 0.1%
2.5%
Show abstract

Background: Global HIV programmes are transitioning virally suppressed children and adolescents with HIV (CAWH) from prior regimens to dolutegravir-based antiretroviral therapy (ART). However, the supporting evidence largely stems from randomised trials in viraemic CAWH. The effect of transition for virally suppressed CAWH is unknown. Methods: We used observational, de-identified data from 724 clinics in KwaZulu-Natal, South Africa. We sequentially emulated three distinct target trials to estimate the effect of transitioning to dolutegravir-based ART in three paediatric populations: i) ages 8-17 years taking efavirenz-based ART, ii) 8-17 years taking ritonavir-boosted lopinavir (LPV/r)-based ART, and iii) 0-7 years taking LPV/r-based ART, all with a last viral load <1,000 copies/mL. The risk difference (RD) of death or viraemia >1,000 copies/mL through 12 and 24 months was estimated using an inverse probability weighting approach. Findings: From January 2020 to August 2024, 37,145 CAWH contributed 454,081 person-trials. In CAWH initially taking efavirenz, the standardised 12-month risk of death or viraemia was 11.9% with continued efavirenz and 6.7% with transition to dolutegravir (RD -5.2 [95% CI -5.8 to -4.6]). In older CAWH initially taking LPV/r, these risks were 17.8% and 9.5%, respectively (RD -8.3 [-10.0 to -6.8]). In younger children, the respective risks were 15.8% and 6.7% (RD -9.0% [-12.7 to -5.4]). Where available, 24-month endpoints showed slightly greater RDs. Interpretation: This large-scale, causal analysis highlights improvements in viral suppression and strongly supports ongoing transition to dolutegravir-based ART for virally suppressed CAWH. Funding: Gates Foundation, National Institute for Health and Care Research, Swiss National Science Foundation

11
Bedside execution, not schedule mismatch: characterizing inpatient carbidopa-levodopa administration timing in Parkinson disease

Plagenz, J.; Lin, A.; Harlow, T.

2026-08-18 health systems and quality improvement 10.64898/2026.08.16.26360535 medRxiv
Top 0.1%
2.5%
Show abstract

Background: Timely carbidopa-levodopa administration is a recognized inpatient safety priority in Parkinson disease, and mistiming is common, but where in the medication-use process it arises is uncharacterized. Objectives: To localize where inpatient mistiming arises and where to target intervention. Methods: In a single-center retrospective analysis of hospitalized adults with Parkinson disease on home carbidopa-levodopa, each dose's administration time was compared with the individualized home schedule. Mistiming was defined a priori as more than 15 minutes from the home time (Parkinson's Foundation Hospital Care Standard 2). We characterized the deviation distribution, tested whether administrations tracked the schedule or the standard grid, and examined length-of-stay and readmission. Results: Across 947 doses in 101 patients, ordering was accurate, yet 62.9% (596 of 947) missed the home time by more than 15 minutes and 99% of patients had at least one mistimed dose. Administrations tracked the individualized schedule almost exactly (Pearson r 0.98), not the standard grid: only 10% fell within 15 minutes of the default times, and the median dose sat 24 minutes from its home time but 76 from the nearest default. Deviation was symmetric drift (median absolute deviation 24 minutes; 16.5% beyond 60 minutes). Conclusions: Mistiming in this study reflected imprecise bedside execution, not ordering or a mismatch between fixed rounds and individualized regimens. These findings may point medication-safety efforts toward protecting bedside administration as complementary redesigning orders.

12
The Heartbeat Study: Feasibility and Advertisement Costs of Implementing a Digital Strategy to Enhance Diversity in the LIBREXIA-AF Clinical Trial

Hussain, T.; Wang, Y.; Chen, Y. Q.; Olson, G.; Panitch, B.; Clemins, K.; Elkarra, N.; Lhamo, K.; Odenwald, N.; Hufner, D.; Jain, S.; Quall, M.; Anderson, C.; Perez, M. V.

2026-08-28 cardiovascular medicine 10.64898/2026.08.24.26361277 medRxiv
Top 0.1%
2.1%
Show abstract

Background: Recruitment of diverse participants remains a challenge in cardiovascular clinical trials. Little is known about how recruitment efficiency and advertising costs with web-based tools vary across US communities. We evaluated an online recruitment platform and examined the cost of acquiring both all-comers and diverse participants in relation to community-level income. Methods: The Heartbeat Study evaluated a digital recruitment strategy to identify US participants for the ongoing Phase 3 LIBREXIA-AF trial. Online advertisements directed individuals with atrial fibrillation to a pre-screening website, where demographic and health data were collected. Advertising impressions, clicks, and costs were recorded. Participant ZIP codes were linked to Core Based Statistical Areas (CBSAs) and CBSA-level income. We measured recruits from underrepresented groups (women, African Americans, Latinos) completing online registration per $100,000 in advertising expenses. Click-weighted linear regression evaluated associations between CBSA income and advertising efficiency. Results: A total of 1,406 recruits completed online registration, with 1,319 participants from 260 CBSAs included in the geographic analysis. Participants were 73 years old on average; 547 (41.5%) were women, 59 (4.5%) African American, and 44 (3.3%) Latino. A total of $163,949.13 was spent on 82,681,711 impressions and 454,750 clicks. Recruits per $100,000 in advertising spend were 334 for women, 36 for African Americans, and 27 for Latinos. CBSA-level income was modestly inversely associated with cost per impression (R2=0.058; p<0.001) and cost per click (R2=0.038; p=0.005), but not recruitment yield for African Americans (p=0.99), Latinos (p=0.37), or women (p=0.21) (R2 range, 0.000-0.13). Conclusions: In this national analysis, online advertising enabled broad engagement across diverse US communities, but income was not associated with recruitment yield among women, African American, or Latino participants. Minority representation remained limited, suggesting digital recruitment alone may be insufficient to improve trial diversity. Targeted, culturally and linguistically tailored strategies may be needed to enhance diverse recruitment.

13
Increasing Lung Cancer Screening Participation Using an Informational Video Nudge: A Randomized Feasibility Trial

Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.

2026-09-01 health systems and quality improvement 10.64898/2026.08.28.26361654 medRxiv
Top 0.1%
2.0%
Show abstract

Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.

14
Vaccination of people with HIV with BG505 SOSIP.v4.1-GT1.1: An interim safety analysis of the investigator-initiated RENEW-SHCS Phase I trial

Poulose, R.; Kusejko, K.; Eichenberger, A.; Manrique, A.; Nemeth, J.; Braun, D. L.; Caringi, I. C.; Mahomed, S.; Garrett, N.; Aceto, L.; Kovari, H.; Huber, M.; Schanz, M.; Kouyos, R. D.; Caskey, M.; Sanders, R. W.; Moore, P. W.; Rauch, A.; Guenthard, H. F.; Trkola, A.

2026-08-27 hiv aids 10.64898/2026.08.24.26360985 medRxiv
Top 0.1%
1.9%
Show abstract

Background: Vaccination of people with HIV (PWH) on suppressive antiretroviral therapy (ART) represents a novel approach for evaluating candidate broadly neutralizing antibody (bnAb) immunogens for preventive and therapeutic HIV vaccines. Given pre-existing immunity in PWH, the safety of this approach requires careful assessment prior to broader application. Here, we report on the design and safety of the RENEW-SHCS study which evaluates the immunization of PWH with BG505 SOSIP.v4.1-GT1.1, an immunogen engineered to induce precursors of CD4 binding site (CD4bs)- and V2-apex targeting bnAbs. Methods. RENEW-SHCS is a phase I, open-label, non-randomized vaccination trial evaluating a single dose of the recombinant germline-targeting envelope trimer BG505 SOSIP.v4.1-GT1.1 (GT1.1), adjuvanted with 3M052-AF and Aluminum hydroxide (alum), in PWH on suppressive ART enrolled from the Swiss HIV Cohort Study. Participants were previously classified as bnAb or non-neutralizing antibody (nnAb) inducers, with a target enrollment of 15 per group, and were monitored for safety and immunogenicity for 24 weeks while continuing standard ART. Due to an out-of-specification stability measurement of adjuvant 3M052-AF the trial was paused after 23 immunizations and subjected to an unscheduled interim safety and reactogenicity assessment comprising protocol defined outcome measures (adverse events, clinical laboratory measurements and HIV-1 viral load). Results. Twenty-three participants (10 bnAb and 13 nnAb inducers, median age 59 years, 17 male / 6 female) were vaccinated between March and August 2025 before interruption of the trial. All participants completed follow-up with full protocol adherence. The interim-safety analysis confirmed that no vaccine-related serious adverse events occurred. Solicited local (96%) and systemic (83%) reactions were common, predominantly grade 1-2, transient, and self-limited. Transient laboratory changes occurred but mostly remained within the normal range, with no vaccine-related grade 3 abnormalities. We observed predominantly transient local and systemic reactions, which were similar or milder to the reactogenicity profile reported for immunization of adult people without HIV (PWOH) with GT1.1 adjuvanted with AS01b reported in the IAVI C101 trial. No viral rebound under ART occurred. One participant experienced two viral blips (>50 HIV-1 RNA copies/ml), one before and one 16 weeks after vaccination with subsequent re-suppression. All others maintained viral suppression (<50 copies/ml) throughout follow-up. Conclusion. RENEW-SHCS demonstrated a favorable safety and reactogenicity profile of single dose immunization with GT1.1 in PWH, comparable to that observed in PWOH. The findings of this phase I study support the feasibility of vaccinating ART-treated PWH in trials of preventive and therapeutic HIV vaccine strategies.

15
Performance of an Ambient Generative AI Documentation Tool in a Linguistically Diverse Clinical Setting

Aldis, R.; Wang, S.; Sage, M.; Metzmaker, M.; Galvin, H.

2026-08-17 health systems and quality improvement 10.64898/2026.08.14.26360467 medRxiv
Top 0.1%
1.8%
Show abstract

Ambient artificial intelligence scribes are being increasingly used in healthcare to improve efficiency and reduce provider clinical documentation burden, yet their performance across linguistically diverse patient populations is not well characterized. We conducted a retrospective analysis of 54,160 outpatient encounters within a U.S. safety net health system to evaluate the performance of an artificial intelligence documentation tool in English and non-English clinical encounters, and in encounters where an interpreter or bilingual provider was present. Documentation performance was measured by the percentage of words in the final note that were generated by the ambient AI documentation tool and not edited by the provider. Associations between language factors and documentation performance were measured using Generalized Estimating Equations with exchangeable correlation structures to account for clustering of multiple encounters within unique patients. Univariable models were fitted to estimate the odds of adequate performance by language and interpreter modality, and a multivariable interaction model was used to evaluate within-language differences between bilingual providers and interpreter-mediated encounters. Non-English encounters were 21% to 25% less likely than English encounters to achieve the same performance threshold. There was no significant difference in generative documentation performance between interpreter-mediated and bilingual provider encounters. These findings underscore the importance of equity-focused evaluation and multilingual model refinement to ensure that artificial intelligence documentation benefits are distributed fairly across diverse patient populations.

16
Systematic Data Fitness Assessment Improves Validity and Replicability of Research Using Real-World Data

Razzaghi, H.; Wieand, K.; Pinkney, A.; Bailey, C.

2026-08-10 epidemiology 10.64898/2026.08.05.26359818 medRxiv
Top 0.1%
1.7%
Show abstract

Research replication is essential to build trust in evidence produced from real-world data. However, methods for conducting and reporting these studies are lacking, particularly related to data quality and fitness assessments. We replicated a single-center study from Children's Hospital of Atlanta in a multi-institutional learning network (PEDSnet) to evaluate the long-term effects of hydroxyurea in children with severe sickle cell disease (SS/S{beta}0 genotype). An AS-IS arm applied the original study's criteria with no major data quality adjustments, while a Data Fitness Enhanced (DFE) arm used systematic data fitness assessment to inform adjustments to cohort inclusion criteria and variable definitions; both arms then replicated the original study's primary analyses. Data quality checks in the DFE arm refined cohort criteria and improved hydroxyurea capture, drug era computation, and hematology specialist mapping. The DFE cohort produced average treatment effects with higher face validity and greater concordance with the original study (e.g., change in ED visits: -0.44 (CI -0.60, -0.26) versus -0.36 (CI -0.57, -0.16) in the original study) than the AS-IS cohort (-0.08 (CI -0.26, 0.09)), which yielded several implausible results. These findings show that superficially plausible cohort characteristics do not guarantee valid results without transparent, systematic data fitness assessment.

17
Quality, consistency, and clinical safety of AI-generated versus clinician-written clinical notes: a multi-country paired simulation study

Bergman, H. I.; Liu, V.; Austin, B.; Ali, S.; Fiedler, M.; Sandiford, C.; Blanchard, R.; Casanovas, C. L.; Pedrazzini, G.; Markopouliotis, T.; Vermersch, F.

2026-08-21 health informatics 10.64898/2026.08.18.26360701 medRxiv
Top 0.1%
1.7%
Show abstract

Background Ambient AI documentation tools, known as scribes, are entering routine clinical practice at scale, but the evidence comparing the notes they produce against clinician-written notes is dominated by single-site, single-language studies that rely on human review to find errors, a method known to miss most documentation errors. Methods We conducted a paired simulation across five countries and languages (Cambridge/English, Barcelona/Spanish, Milan/Italian, Paris/French, Cologne/German; 385 paired consultations, 770 notes). From each actor-performed consultation, an AI scribe (Heidi) and a junior-to-middle-grade clinician independently produced a note. Notes were scored on the PDQI-9 by evaluators blinded to authorship. Documentation errors were identified by two methods of deliberately different sensitivity - clinician adjudication, and a calibrated automated reviewer externally validated against a blinded ten-clinician panel - then graded for clinical risk by a three-model panel. The co-primary outcomes were PDQI-9 total and Critical+High error burden, the latter reported under both detection arms. The analysis plan was registered before any pooling across sites. Results AI notes scored higher than clinician notes on the PDQI-9 (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; Cohen dz=0.55), consistently across all five sites (dz 0.41-0.75), and were less dispersed (5.7% of AI vs 27.8% of clinician notes fell below the study pre-specified low-score threshold (<32)). On the principal safety outcome - the paired probability that a note carried [&ge;]Critical+High error - clinician notes were affected more often under both detection arms: 61.0% versus 24.4% by the calibrated reviewer (relative risk 2.50, 95% CI 2.09-3.00) and 21.8% versus 6.2% by clinician adjudication (relative risk 3.50, 95% CI 2.32-5.27). The difference was largest for omissions. Unaided clinician review identified roughly 12% of the errors the calibrated reviewer retained, and a smaller fraction in AI notes than in clinician notes. Conclusions In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians. The magnitude of the safety difference depends on the sensitivity of error detection, so we report both detection regimes and bound rather than point-estimate the absolute error rate. Extension to live practice, consultant-authored documentation, and notes as filed after clinician editing remains to be established.

18
Drivers of Oncologist Preference of AI-Generated Literature Review in a Randomized Mixed-Methods Study

Bunning, B. J.; Weng, Y.; Wu, D. J.; Hui, G.; Hope, J. E.; Pandurangan, V.; Lopez, I.; Everett, S.; Chen, J. H.; Desai, M.

2026-08-27 health informatics 10.64898/2026.08.24.26361252 medRxiv
Top 0.1%
1.7%
Show abstract

Doctors increasingly rely on AI in the clinic, yet which report features make AI-generated responses useful and trustworthy remains unclear. In this randomized mixed-methods study, 34 oncology physicians provided 294 ratings of four blinded AI systems across five vignettes, alongside 20 semi-structured interviews analyzed with a prespecified LLM-assisted qualitative pipeline. Despite similar references, an evidence-graded report adapted from OpenEvidence was rated significantly lower in overall utility than standard OpenEvidence (mean difference, -0.96; 95% CI, -1.26 to -0.66; P<.001). Qualitative analysis identified six themes and seven design requirements. Oncologists valued rapid orientation, evidence retrieval, and verification, preferring concise, scannable reports with quantitative outcomes, recognizable bolded guidelines, explicit uncertainty, and verifiable citations. Trust deteriorated with citation mismatch, buried provenance, evidence misclassification, overconfident recommendations, and poor organization. Evidence presented differently can alter perceptions of clinical utility and trust; accuracy alone is insufficient, and report design must also be empirically evaluated.

19
Engaging Zimbabwean men and stakeholders in the co-adaptation of peer-delivered HIV self-testing: iterative prototyping of the IMPERATIVE Trial

McGowan, M.; Maswera, R.; Chisvo, L.; Moorhouse, L.; Dzamatira, F.; Mandizvidza, P.; Tsenesa, B.; Otambo, W.; Inghels, M.; Harling, G.; Mee, P.; Baernighausen, T.; Gregson, S.; Nyamukapa, C.; Tanser, F.; Skovdal, M.

2026-08-11 hiv aids 10.64898/2026.08.10.26360082 medRxiv
Top 0.1%
1.7%
Show abstract

Introduction: HIV testing and pre-exposure prophylaxis (PrEP) are efficacious HIV prevention strategies, but uptake remains low among Sub-Saharan African men. Peer-delivered approaches may improve engagement. We developed an intervention combining peer-delivered oral HIV self-testing (HIVST) with incentivized peer referral to HIV services and an SMS-based HIV risk assessment among men in eastern Zimbabwe (IMPERATIVE Trial: NCT06370923). We co-adapted the intervention through iterative prototyping (IP) to enhance its acceptability, feasibility, and potential effectiveness. Methods: From November 2023 to June 2024, we implemented a novel IP framework to refine and test the intervention. Four primary distributors (PDs) were trained to deliver HIVSTs to three peers and refer them to clinic services. Peers could become secondary distributors (SDs), obtain HIVSTs from community hubs and distribute them further. Qualitative data were collected alongside intervention testing to adapt the intervention over two iterations. Activities included three forum theatre workshops, one community advisory board meeting, 25 in-depth interviews, four focus group discussions, and eight observational reports involving men, implementers, stakeholders, and advisory board members. Additionally, 20 men completed baseline and one-week follow-up surveys. Quantitative data were analysed descriptively; qualitative data were analysed using thematic analysis. Results: During testing, HIVST uptake was 100% among PDs, 90% among PD-recruited peers, and 63% among SD-recruited peers. Among self-testers, 50% sought confirmatory testing and about one-quarter initiated PrEP (PDs 25%, PD-recruited peers 30%, SD-recruited peers 25%). Participants viewed the intervention positively and anticipated increased HIV testing and PrEP initiation. Four areas for refinement were identified: recruitment, information dissemination, incentives, and socio-cultural factors. Participant recommendations were adopted before randomised controlled trial testing. Conclusion: Peer-delivered HIVST with referral to HIV services shows promise for engaging Zimbabwean men. The IP framework incorporating participant recommendations enhanced intervention design and delivery within the IMPERATIVE trial. This methodology may inform future intervention development in similar settings.

20
TrialCode Agent: LLM-Assisted Clinical Code-Set Construction for Trial Emulation

Habibdoust, A.; Sajjad, A.; Hernandez, D.; Patel, K.; Song, X.

2026-08-23 health informatics 10.64898/2026.08.20.26360962 medRxiv
Top 0.1%
1.7%
Show abstract

Objective Translating free-text clinical trial criteria into computable code sets is a valuable standardization practice that is necessary for producing reproducible real-world evidence studies but requires standardized interpretation across multiple clinical vocabularies. Methods We developed TrialCode Agent, a hybrid-large language model (LLM)-terminology verification agent that generates, formats, verifies, and expands candidate codes from free-text clinical criteria. The system supports ICD-9-CM diagnoses and procedures, ICD-10-CM, ICD-10-PCS, LOINC, and RxNorm medication concepts. We compared Baseline, Hybrid biomedical retrieval-augmented generation (RAG), and terminology-guided Family expansion pipelines using Claude, GPT Qwen, and MedGemma on 40 criteria from 11 trial groups. Performance was evaluated against expert-built reference code sets using exact-code precision, recall, and F1. Results The optimal pipeline varied by model. Claude with Baseline achieved the highest performance (precision 0.755, recall 0.619, F1 0.680), followed by GPT-5.5 with Baseline (precision 0.569, recall 0.658, F1 0.610), Qwen with Hybrid biomedical RAG (precision 0.656, recall 0.470, F1 0.548), and MedGemma with Family expansion (precision 0.487, recall 0.316, F1 0.383). Hybrid biomedical RAG improved aggregate F1 only for Qwen but increased GPT-5.5 RxNorm F1 from 0.320 to 0.909. Macro-averaged results showed criterion-level gains despite lower micro-averaged aggregate performance. Family expansion increased recall across models but generally reduced precision. In staged verifier ablation, micro-F1 increased from 0.254 before verification to 0.505 after final verification and expansion. Existence/vocabulary checking removed 2,594 false-positive codes, and acceptance filtering removed 952 additional false-positive codes before controlled expansion. Conclusions Combining LLM-based clinical interpretation with deterministic terminology verification produces auditable, database-ready code sets, but retrieval and broad family expansion do not consistently improve exact-code performance. Retrieval was particularly useful for RxNorm mapping, whereas overly broad or incomplete candidate generation remained the main source of error. Deterministic verification improves code validity and query readiness but cannot replace accurate clinical interpretation.